Tag
45 articles
Alibaba's Qwen-Image-3.0 introduces advanced image generation capabilities, including support for 4,500-token prompts, readable ten-pixel text, and complex layout rendering in a single pass.
This explainer explores Alibaba's Qwen 3.8, a multimodal AI model with 2.4 trillion parameters that rivals top-tier models like Fable 5. We examine its architecture, training methods, and implications for the future of large language models.
Thinking Machines Lab launches Inkling, a 975-billion-parameter open source model trained to understand video and audio, positioning itself against competitors like Anthropic and OpenAI.
Learn how to set up and run inference with the Inkling multimodal AI model from Thinking Machines Lab, including text and image processing with controllable thinking effort.
Researchers have reconstructed the VideoAgent workflow into a functional, API-key-free multi-agent system for AI-powered video editing, enabling natural language interactions and automated video processing.
Meta Superintelligence Labs introduces Muse Spark 1.1, a multimodal reasoning model for agentic tasks, featuring a 1,000,000-token context window and multi-agent delegation capabilities.
The next leap in AI video is not just about improving visual fidelity, but teaching avatars to see, hear, and interact in real time. This shift is transforming how we think about digital experiences.
Explains Apple's advanced 'Apple Intelligence' framework, detailing how transformer-based architectures, multimodal processing, and privacy-preserving techniques will revolutionize AI assistants and human-computer interaction.
Google AI announced major advancements in multimodal models, safety measures, and enterprise applications in May 2026. The company's Gemini 2.0 release represents a significant leap in AI capabilities and accessibility.
Google Deepmind's Gemma 4 12B is an open-source multimodal AI model that runs efficiently on laptops with just 16 GB of RAM, nearly matching the performance of its larger 26B counterpart.
Alibaba's Qwen team launches Qwen3.7-Plus, a multimodal AI model on the Bailian platform, featuring vision understanding, deep reasoning, tool invocation, and autonomous iteration.
Chinese AI company MiniMax has unveiled M3, the first open-weight model combining top-tier coding performance, a one-million-token context window, and native multimodality, challenging proprietary leaders in the AI space.